certainty equivalent
Risk-Sensitive Q-Learning in Continuous Time with Application to Dynamic Portfolio Selection
This paper studies the problem of risk-sensitive reinforcement learning (RSRL) in continuous time, where the environment is characterized by a controllable stochastic differential equation (SDE) and the objective is a potentially nonlinear functional of cumulative rewards. We prove that when the functional is an optimized certainty equivalent (OCE), the optimal policy is Markovian with respect to an augmented environment. We also propose \textit{CT-RS-q}, a risk-sensitive q-learning algorithm based on a novel martingale characterization approach. Finally, we run a simulation study on a dynamic portfolio selection problem and illustrate the effectiveness of our algorithm.
A Appendix of Proofs 1 A.1 Proof of Thm.3.2
Eqn. 3) is equivalent to optimizing CL (InfoNCE, cf. To complete the proof, we start with giving some important notations and theorem. Here we simply disregard the constant term present in Eqn. 4 as it does not impact optimization, and From the Thm.3.2, we have the equivalence between InfoNCE and CL-DRO. By using McDiarmid's inequality in Thm A.4,for any ϵ, we have: While Corollary 3.4 has already been proven in [ We start with introducing a useful lemma. Then the CL-DRO objective is the tight variational estimation of ϕ -divergence.
How Personality Traits Shape LLM Risk-Taking Behaviour
Hartley, John, Hamill, Conor, Batra, Devesh, Seddon, Dale, Okhrati, Ramin, Khraishi, Raad
Large Language Models (LLMs) are increasingly deployed as autonomous agents, necessitating a deeper understanding of their decision-making behaviour under risk. This study investigates the relationship between LLMs' personality traits and risk propensity, employing cumulative prospect theory (CPT) and the Big Five personality framework. We focus on GPT-4o, comparing its behaviour to human baselines and earlier models. Our findings reveal that GPT-4o exhibits higher Conscientiousness and Agreeableness traits compared to human averages, while functioning as a risk-neutral rational agent in prospect selection. Interventions on GPT-4o's Big Five traits, particularly Openness, significantly influence its risk propensity, mirroring patterns observed in human studies. Notably, Openness emerges as the most influential factor in GPT-4o's risk propensity, aligning with human findings. In contrast, legacy models like GPT-4-Turbo demonstrate inconsistent generalization of the personality-risk relationship. This research advances our understanding of LLM behaviour under risk and elucidates the potential and limitations of personality-based interventions in shaping LLM decision-making. Our findings have implications for the development of more robust and predictable AI systems such as financial modelling.